Back

Journal of Bioinformatics and Systems Biology

Fortune Journals

Preprints posted in the last 30 days, ranked by how well they match Journal of Bioinformatics and Systems Biology's content profile, based on 15 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Kiosc: an integrated platform for managing bioinformatics data analysis containers

Marotta, F.; Stolpe, O.; Obermayer, B.; Weiner, J.; Holtgrewe, M.; Beule, D.; Nieminen, M.

2026-08-24 bioinformatics 10.64898/2026.08.20.745983 medRxiv
Top 0.1%
1.4%
Show abstract

In many bioinformatic data analysis projects, it is convenient to visualize plots and results through an interactive web app or dashboard. These interactive reports can then be shared with customers, collaborators, or the general public. Publishing and sharing these apps is not straightforward, becoming especially cumbersome when the number of projects and customers start growing. Docker containers offer a convenient way to package, distribute, and run interactive web apps, and their use is already widespread in the bioinformatics community. We developed Kiosc to simplify the orchestration of containerized web apps, organize them into projects, and regulate access control. We implemented it as a web server based on the Django framework, with a user- and admin-friendly interface as well as a REST API for programmatic tasks. Users can select Docker containers packaging apps like Plotly Dash, Shiny, or Quarto, and configure them to display the results of their analysis. Kiosc runs the containers with the appropriate network configuration and acts as a proxy to the web services running inside the containers. We have been maintaining a Kiosc instance for more than 5 years, serving 321 containers in 150 projects across multiple institutions. In this article, we introduce the main functionality in Kiosc and describe four use-cases that show how Kiosc can prove helpful to the broader bioinformatics community, such as configuring and running web apps for the interactive visualization of workflow results, and publishing companion apps for scientific articles. Kiosc is a self-hosted platform for publishing web apps, which doesn't require significant expertise in either Docker or network administration to be deployed. It provides a similar service to Kubernetes, but with a convenient web interface and much lower administration overhead.

2
Rapid PCR-based screening system for detection of type II CRISPR-Cas loci in bacterial species

Bibi, A.; Iqbal, T.; Ilyas, K.; Nosheen, A.

2026-08-26 molecular biology 10.64898/2026.08.24.746701 medRxiv
Top 0.1%
1.1%
Show abstract

The Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) and associated nuclease gene (Cas), originating from the bacteria acquired immune system, have revolutionized gene editing technology. In this regard, type II (Cas9) been extensively studied and widely applied CRISPR system so far. The mechanism for precise manipulation of genomic sequences is guided by small RNA called CRISPR RNA (crRNA). In this study we devised and optimized CRISPR-Cas9 screening system based on Cas9 gene detection, targeting a conserved part of recognition domain (REC) consisting of arginine rich bridge helix (BH). We used hemi-nested PCR approach for screening sensitivity and reproducibility. The recombinant E. coli DH5 alpha containing the pRGEB32 vector (DH5 alpha/pRGEB32) with the Cas9 gene was used for system optimization. Subsequently, the screening system was applied and validated on different environmental bacterial strains including Alcaligenes faecalis and Pseudomonas stutzeri, isolated from sewerage samples. The optimized hemi-nested PCR resulted in amplification of targeted region in environmental bacterial strains and results were reproduced successfully. Furthermore, nucleotides and amino acid sequence, motif and domain analysis of PCR products, confirmed the targeted Cas9 REC-BH domain. Presently, no rapid and cost effective CRISPR-Cas screening system is available except expensive whole genome sequencing approach. Our investigation aimed to device rapid and cost effective screening system for identification of new variants of Cas9 proteins in environmental bacterial species. In this context, the developed Cas9 gene-based CRISPR-Cas screening system (C9CSS) may be a potential rapid screening tool to identify new Cas9 orthologs in different bacterial genomes with improved functions.

3
Ori-Finder-Arch: An Updated Web Server for the Annotation and Visualization of Archaeal Replication Origins

You, Z.; Zhang, Z.; Luo, H.; Gao, F.

2026-08-19 bioinformatics 10.64898/2026.08.15.744077 medRxiv
Top 0.2%
0.9%
Show abstract

Archaea are promising chassis organisms in biotechnology, and the accurate annotation of their chromosomal replication origins (oriCs) is the key to unlocking their full potential. However, the existing Ori-Finder 2 web server suffers from low accuracy, slow speed, and limited scalability. In this study, we present Ori-Finder-Arch, an updated web server for high-performance oriC prediction in archaea. This pipeline integrates HMMER-based replication initiation protein (RIP) annotation, refined consensus motif recognition, and GC profile-based DNA unwinding element (DUE) detection. On a benchmark set of experimentally validated oriCs, Ori-Finder-Arch achieved a recall of 95.6% and a precision of 86.0%, substantially outperforming Ori-Finder 2 (62.2% and 63.6%, respectively), while running 4.75 times faster and supporting diverse assembly levels. When applied to the available archaeal assemblies, it successfully annotated 17,472 oriCs. Meanwhile, the web server provides interactive visualizations at different levels. In conclusion, Ori-Finder-Arch offers an efficient, accurate, and user-friendly platform for advanced studies of archaeal DNA replication initiation and synthetic biology applications, and is freely available at https://tubic.org/Ori-Finder-Arch/ and https://tubic.tju.edu.cn/Ori-Finder-Arch/.

4
Alcama expressed in blood retina barrier and Muller glia is involved in zebrafish retina regeneration

Thomas Michael, S.; Allan, K.; Rini, M.; DiCicco, R.; Ramos, M.; Yuan, A.

2026-08-25 cell biology 10.64898/2026.08.24.746827 medRxiv
Top 0.2%
0.9%
Show abstract

Activated leukocyte cell adhesion molecule A (Alcama) plays a role in axonal guidance, cell differentiation, and retinal lamination in a developing retina and was identified as a marker for activated Muller glial cells in adult zebrafish. However, its spatiotemporal localization and its involvement in retina regeneration remains unclear. Here we induced focal photoreceptor damage in zebrafish using laser photocoagulation and examined the expression and localization of Alcama at different time points post lesion. Immunohistochemistry in wild type fish and Tg(kdrl-EGFP) fish showed Alcama localized to the blood retina barrier with increased expression in Muller glial end feet and radial processes in a regenerating retina. To confirm its role in retina regeneration, alcama expression was transiently knocked down using morpholinos in adult fish. Scanning laser ophthalmoscopy, Zpr1 immunostaining and EdU staining showed delayed retina regeneration in alcama knockdown fish, indicating a possible role for Alcama in zebrafish retina regeneration.

5
SVlog: a logic programming framework for understanding structural variation in genomic disease

Gudkov, M.; Reis, A. L. M.; Kumaheri, M.; Deveson, I. W.

2026-08-21 bioinformatics 10.64898/2026.08.11.744322 medRxiv
Top 0.3%
0.8%
Show abstract

Structural variants (SVs) are a diverse group of genetic variants defined by a minimum size of 50 base pairs. SVs account for the majority of all variant bases in a persons genome and are commonly implicated in inherited disease and cancer. However, SV analysis is complex due to their wide variation in type and size, degree of polymorphism, involvement of repetitive sequences, and the myriad ways they may elicit a functional impact, as well as technical factors like imprecise breakpoint detection, and alternative representations of the same event. Despite recent advances in the detection and characterisation of SVs, it remains difficult to assess them beyond basic annotations and comparisons. Here we introduce SVlog, a transparent and extensible meta-programming framework for SV analysis. With the logic programming language Souffle as its engine, SVlog provides a declarative ontology describing relationships among SVs, genes and other genomic elements. Genome annotations and SV datasets - both user-provided and public reference data - are converted into relational facts, to which SVlog applies logical rules that define predicates. Predicates are specific, transparent and deterministic, yet fully flexible and composable, enabling detailed evaluation of SVs without relying on stochastic "black box" approaches. To showcase SVlog, we have developed a ready-made predicate library for SV annotation, comparison and prioritisation in the context of rare inherited disease. Despite its compact codebase, SVlog evaluates more than 50 input predicates to generate over 70 informative output predicates. It synthesises evidence from population and clinical genomic databases, and applies a tiered filtering strategy to identify candidate pathogenic SVs in patients with inherited disease. By focusing on explainability and modularity, SVlog offers a fast, reliable library for SV analysis and is a powerful deterministic alternative to traditional bioinformatics pipelines for clinical variant curation.

6
Abundant Glomerular Neutrophil Extracellular Traps in C3 Glomerulopathy

O'Sullivan, K.; khandelwal, p.; Walker, P. D.; hickey, m.; Licht, C.

2026-09-01 immunology 10.64898/2026.08.27.747386 medRxiv
Top 0.3%
0.6%
Show abstract

Introduction: C3 glomerulopathy (C3G) is driven by fluid-phase alternative complement pathway dysregulation, with emerging evidence linking glomerular neutrophil infiltration to disease severity. Neutrophil extracellular traps (NETs) are implicated in other forms of glomerulonephritis. However, their participation in the pathogenesis of C3G remains undefined. Methods: Kidney biopsies from 33 patients with C3G (15 with dense deposit disease [DDD] and 18 with C3 glomerulonephritis [C3GN]) were compared with 15 anti-neutrophil cytoplasmic antibody associated vasculitis (AAV) biopsies as a neutrophil-rich disease control in this retrospective cross-sectional study. Glomerular neutrophils and NETs were identified using immunofluorescence, staining for myeloperoxidase, citrullinated histone H3, peptidyl arginine deiminase-4, and DNA. Supervised machine learning was used to quantify glomerular NET formation, and the data were correlated with kidney function at time of biopsy using linear regression. Results: Intraglomerular NETs were abundant and detected in the majority of glomeruli in C3G biopsies. Compared with AAV, C3G showed a significantly higher fraction of neutrophils forming NETs, despite similar neutrophil counts per glomerulus. NET abundance was similar in DDD and C3GN. In exploratory analyses, a greater proportion of glomeruli containing NETs was associated with lower kidney function (estimated glomerular filtration rate) at biopsy, and this association remained significant after adjustment for age, C3G subtype, and interstitial fibrosis. Conclusions: These observations demonstrate that intraglomerular NETs are a common and prominent observation in C3G and are associated with reduced kidney function at biopsy. These findings raise the possibility that NET deposition in glomeruli is a previously unrecognized driver of glomerular injury in C3G.

7
Gene model for the ortholog of Ilp3 in Drosophila pseudoobscura

Lieser, B. C.; Laskowski, L. F.; Huber, R.; Kolker, K. O.; Arsham, A. M.; Rele, C. P.; Toering Peters, S.

2026-08-23 genomics 10.64898/2026.08.19.745830 medRxiv
Top 0.3%
0.6%
Show abstract

Gene model for the ortholog of Insulin-like peptide 3 (Ilp3) in the D. pseudoobscura Apr. 2013 (BCM-HGSC Dpse_3.0/DpseGB3) Genome Assembly (GenBank Accession: GCA_000001765.2) of Drosophila pseudoobscura. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

8
Click-Prep: An Interactive Data Preparation Tool for Click-qPCR

Kubota, A.; Tajima, A.

2026-08-24 bioinformatics 10.64898/2026.08.20.745930 medRxiv
Top 0.4%
0.6%
Show abstract

Click-qPCR is a browser-based application for relative qPCR analysis that requires a tidy-format CSV file containing four columns: sample, group, gene, and Cq. Preparing this input from qPCR instrument output typically requires manual reformatting and calculation of mean Cq values for technical replicates. To simplify this process, we developed Click-Prep (https://kubo-azu.shinyapps.io/Click-Prep/), an interactive web-based application designed specifically to create Click-qPCR input files. Click-Prep imports CSV, TXT, TSV, and XLS/XLSX files and supports skipping of instrument-generated metadata rows, interactive column mapping, and manual assignment of experimental groups. Users can review technical-replicate measurements, exclude selected rows according to predefined quality-control criteria, and calculate mean Cq values for each sample-group-target combination. Missing or nonnumeric Cq values are flagged for review and must be resolved before the mean is calculated. Click-Prep can also combine compatible formatted CSV files, such as datasets obtained from separate qPCR plates. The resulting dataset is exported as a standardized CSV file containing the four fields required by Click-qPCR. By integrating these operations into a guided browser-based workflow, Click-Prep enables users to prepare Click-qPCR input files rapidly and consistently without programming.

9
CyChat: a conversational Cytoscape app for no-code, reproducible network analysis

Liebold, J.; Stahl, M.; Schulze, J.-O.; Razavi, M. M.; Bader, G. B.; Kurtz, S.; Baumbach, J.

2026-09-01 bioinformatics 10.64898/2026.08.28.747833 medRxiv
Top 0.4%
0.6%
Show abstract

Network-based analyses of molecular interactions are useful for interpreting high-throughput omics data and identifying therapeutic targets. Cytoscape is the standard platform for these tasks, but users face a trade-off between accessible graphical workflows that are difficult to document and reproducible automation in Python or R that requires programming expertise. General-purpose coding assistants can generate Cytoscape Automation scripts, but remain external to Cytoscape. We present CyChat, a Cytoscape Desktop app that integrates a chat interface and a large language model (LLM) agent into the application. CyChat translates natural language into executable Cytoscape Automation workflows, runs generated Python code, and exports chat sessions with executed code as standalone Jupyter notebooks. To reduce setup barriers, CyChat includes an embedded Python runtime and supports both cloud-based and locally hosted LLMs. CyChat was evaluated across ten Cytoscape workflows using seven LLM providers, each represented by one LLM. The strongest configuration achieves a pass rate above 99%. In a qualitative evaluation based on a published network visualization, CyChat completes the task in 1.5-5 minutes, compared with 15-20 minutes for manual GUI workflows by computational biologists. CyChat is available through the Cytoscape App Store at https://apps.cytoscape.org/apps/cychat.

10
Mechanism of Renal Cyst Initiation and Progression Through ETV Transcription Factors and Hedgehog Signaling

Ryu, B.; Ha, L.; Dsouza, D. L.; Boesen, E. I.; Huh, S.-H.

2026-08-26 developmental biology 10.64898/2026.08.21.746191 medRxiv
Top 0.5%
0.5%
Show abstract

Renal cysts are categorized as non-pathogenic simple cysts and pathogenic malignant cysts based on their pathophysiological status. Cyst formation is divided by cyst initiation and cyst progression/promotion. Pathogenic cysts are thought to be developed through continuous initiation followed by progression until pathogenic status is achieved. Although many genetic and environmental factors are identified to cause pathogenic cyst formation, the mechanisms that discriminate cyst initiation and progression are poorly understood. Using genetic mutation models of ETV transcription factors, ETV1, ETV4, and ETV5, and a pharmacological inhibitor of hedgehog signaling, cyclopamine, we identified one of the mechanisms regulating cyst initiation and progression. Nephron specific deletion of ETV4 and ETV5 initiated cyst formation. However, cyst initiation did not continue as animals grow, and a limited number of the initial cysts underwent further growth. Additional deletion of ETV1 was required for continuous initiation in addition to promotion of cyst growth. Furthermore, administration of cyclopamine attenuated promotion of cyst progression but had little effect on cyst initiation. Therefore, we provide evidence that cyst initiation and progression is genetically and molecularly distinct and can be modulated. This information provides new insight into how to control renal cyst initiation and progression and can be used to suppress pathogenic cyst growth.

11
DrosoTracker: a web application with a self-calibrating thermal model for husbandry scheduling and lifespan analysis in Drosophila melanogaster

Asti Tello, G. S.; Melani, M.; Liberman, A. C.

2026-08-11 developmental biology 10.64898/2026.08.10.743933 medRxiv
Top 0.5%
0.5%
Show abstract

Planning husbandry tasks and experiments with Drosophila melanogaster requires converting a target date into development times that depend on the rearing temperature. This calculation needs to be done for each cross, genotype, and temperature, and the risk of error grows quickly. Available laboratory management tools let users register stocks, crosses, and track them, but they do not create schedules based on a clear, adjustable thermal model. To fill that gap, we developed DrosoTracker, a self-contained web application that works offline and predicts Drosophila development with a thermal summation model recalibrated through regression on data from Powsner (1935) (T0 = 11.78 {degrees}C, DD = 116.38 {degrees}C{middle dot}days, R{superscript 2} = 0.997). The model offers an optional two-level calibration driven by user observations. A wild-type strain first adjusts the model to the laboratorys own conditions. Then each genotype is calibrated against that reference using a random-effects shrinkage estimator that accounts for measurement error and between-batch variability. The model creates schedules for husbandry tasks, evaluates adult cohort survival with the Kaplan-Meier estimator and the log-rank test, and calculates sample size for lifespan studies using Schoenfelds formula. The quantitative components were checked against independent references, including Rs survival package and manual calculations. Ongoing work is focused on validating the calibrated model using cohorts specifically bred for this purpose. DrosoTracker runs entirely in the browser, stores data locally, and is available in English and Spanish.

12
Detecting and typing Chlamydia trachomatis strains in metagenomes using the MetaChlam pipeline

Sharma, P.; Dean, D.; Read, T. D.

2026-08-22 bioinformatics 10.64898/2026.08.18.745514 medRxiv
Top 0.5%
0.5%
Show abstract

The Gram negative bacteria Chlamydia trachomatis (Ct), an obligate intracellular human pathogen, is a predominant cause of sexually transmitted infections and ocular trachoma globally, exerting a significant impact on public health. Ct "strains" (major lineages within the species) are known to have different tissue tropisms and be associated with different disease outcomes. Metagenome samples from typical sites where Ct infects (e.g., endocervix, conjunctiva, rectum) rarely contain enough reads for traditional genotyping methods such as Multi-Locus Sequence Typing (MLST) or ompA genotyping. To overcome these limitations, we implemented an ensemble tool called MetaChlam that can accurately classify Ct strains with as few as 250 Ct reads. Using 109 publicly available Ct genomes from naturally circulating strains, we established that an ANI-based threshold of 99.75% was capable of distinguishing Ct strains from each other. We implemented metagenome-based typing using the previously developed LINtax, Strainscan, StrainGE, and Sourmash softwares. MetaChlam integrated the four tools along with custom databases into an automated nextflow pipeline. Using simulated metagenomic reads, we found that our pipeline accurately identified the correct strains in both single strain and multi-strain mixtures of samples. Finally, we showed that MetaChlam had higher specificity for the true presence of Ct reads in NCBI SRA metagenomic datasets than NCBI PebbleScout software. A surprising finding of these analyses was that reads from Ct, an obligate human intracellular pathogen, can be found as contaminants in samples from sites where the organism is almost certainly not present. Overall, our study enhances the characterization and classification of Ct strains and provides protocols for identification and typing of Ct in shotgun metagenome data. The MetaChlam pipeline is available on Github: https://github.com/parul-sharma/MetaChlam.

13
ASAREE: An Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation

Moran, J.; Freda, P. J.; Ghosh, A.; Hernandez, M. E.; Moore, J. H.

2026-08-25 bioinformatics 10.64898/2026.08.20.746074 medRxiv
Top 0.6%
0.5%
Show abstract

Summary: Agentic AI platforms enable the engineering of autonomous workflows but are not designed for experimentation and hypothesis testing. ASAREE (Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation), is an open-source platform to address this gap. ASAREE creates agents, connects to MCP servers and tools, and designs factorial experiments through a visual interface or Python SDK. It records a full provenance trace for every run and routes all model calls through a provider-agnostic bridge that supports local deployments, ensuring data privacy. As a use-case, we use ASAREE to evaluate key design choices in a mutli-agent machine learning pipeline. Across a 2 x 2 x 2 factorial design, more advanced models, greater reasoning effort, and critic agent use significantly increased compute time, token use, cost, and feature count without improving predictive performance. The lowest-cost baseline, Claude Sonnet 5 with medium effort and no critic, achieved the highest mean PR AUC while Claude Opus 5 with extra high effort and a critic agent cost 15.5x more (USD) and ran 13.1x longer while performing worse on average. These findings highlight ASAREE as a robust framework for evaluating agentic system performance and resource efficiency.

14
Serial Immunohistochemistry for High-Dimensional Single-Cell Spatial Analysis of Human Kidney Biopsies

Yang, X.; Marlin, M. C.; Celia, A. I.; Lee, C.-Y.; Cammarata-Mouchtouris, A.; Stephens, T.; Haddad, M.; Bradshaw, L.; Saksena, D.; Buyon, J.; Izmirly, P. M.; Putterman, C.; Kamen, D.; Petri, M.; Accelerating Medicines Partnership: RA/SLE Network, ; James, J. A.; Guthridge, J. M.; Fava, A.; Rosenberg, A. Z.

2026-08-12 pathology 10.64898/2026.08.06.743188 medRxiv
Top 0.6%
0.5%
Show abstract

BackgroundTraditional immunohistochemistry (IHC) with chromogen detection has limited multiplex capacity, detecting at most 4 protein markers per tissue section simultaneously, thereby restricting comprehensive spatial analysis of valuable human biopsies. We developed and validated a robust serial IHC (sIHC) staining method to detect multiple antigens on a single kidney biopsy slide, maximizing data yield for diagnosing and studying complex kidney diseases. MethodsFormalin-fixed, paraffin-embedded kidney biopsy sections were subjected to repeated IHC/imaging cycles with antibody removal using an optimized sodium dodecyl sulfate-glycerol buffer stripping protocol. Images were then co-registered, and analysis was performed using a variety of methodologies, including color deconvolution, cell segmentation, and spatial clustering. ResultsThis optimized sIHC method successfully detected up to 20 antigens on a single slide. Combining image analysis and artificial intelligence software, for example with HALO (Indica Labs), the assay assembles high-dimensional images and enables quantitative histology and single-cell spatial analysis. Using this advanced method, we were able to identify rare cell populations, such as double-negative T cells, that are challenging to detect conventionally. ConclusionWe have developed a validated, high-capacity sIHC protocol that uses standard IHC procedures with commercially available, clinically validated off-the-shelf antibodies. This method is a valuable, cost-effective tool for obtaining extensive, high-dimensional single-cell-resolved spatial data from limited pathology samples, such as a human kidney biopsy.

15
Inhibition of the Notch signaling pathway promotes AQP2 plasma membrane accumulation in renal epithelial cells by depolymerizing actin and reducing endocytosis

Tchakal Mesbahi, A.; Huang, H.; Ross, J. C.; Bouley, R.; Brown, D.

2026-08-12 cell biology 10.64898/2026.08.11.744289 medRxiv
Top 0.6%
0.5%
Show abstract

The Notch signaling pathway plays a central role in development and cell fate determination. Its function depends on tightly regulated intracellular trafficking of the Notch receptor and the Notch intracellular domain (NICD) after cleavage by {gamma}-secretase. Notch signaling is essential for principal cell differentiation within the renal collecting duct and for proximal-distal patterning during kidney development. Notch activity has also been shown to influence the trafficking of several membrane proteins, including nephrin in kidney cells and monocarboxylate transporter 1 in brain endothelial cells. Aquaporin-2 (AQP2) is the key vasopressin-regulated water channel in the collecting duct, and proper AQP2 trafficking and recycling are required for physiologically appropriate urine concentration. To determine whether and, if so, how Notch signaling modulates AQP2 trafficking, we performed studies using LLCPK1 renal epithelial cells stably expressing AQP2 (LLCPK1-AQP2). Exposing cells to 35 M DAPT (which inhibits y-secretase, preventing cleavage and activation of Notch receptor signaling) for 30 min significantly increased AQP2 membrane accumulation in LLCPK1-AQP2 cells as revealed by immunofluorescence staining. Using a rhodamine-transferrin internalization assay, we found that DAPT reduced clathrin-mediated endocytosis by 60%. This blockade increases AQP2 membrane accumulation by preventing the reinternalization of AQP2 that is delivered to the plasma membrane by exocytosis during its constitutive recycling pathway. Using an F-actin polymerization assay, we then found that Notch inhibition decreases F-actin polymerization by de-activating the small GTPase RhoA, using GSTRBD, a substrate that binds to active RhoA, as seen by western blotting using phospho-specific antibodies. Because actin polymerization is required for AQP2 endocytosis, RhoA inhibition by DAPT would result in the decreased internalization of AQP2 that we observed by immunofluorescence. While the mechanism by which DAPT inhibits RhoA activity remains to be determined, our study shows that AQP2 trafficking is regulated by the Notch signaling pathway in vitro and suggests that modulation of Notch signaling may represent a novel strategy to address water balance disorders that involve defects in the AQP2 trafficking process.

16
Efficient Game-Theoretic Explanations for Tree-Based Ensembles via Owen Values

Koh, H.

2026-08-20 bioinformatics 10.64898/2026.08.12.744440 medRxiv
Top 0.6%
0.5%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWShapley-value-based explanations, notably SHAP (SHapley Additive exPlanations), have gained prominence as a principled game-theoretic framework for local explanations and global feature importance. While exact Shapley value computation is exponential in feature count, TreeExplainer exploits the recursive structure of decision trees to achieve polynomial-time computation for tree-based ensembles. In many scientific applications, however, features are naturally organized into a priori groups reflecting domain knowledge, requiring explanations both across and within groups. The Owen value extends the Shapley value through a two-stage allocation rule that incorporates group structure while preserving fairness properties; yet, efficient algorithms for its computation remain limited. In this paper, we propose exact and Monte Carlo algorithms for computing Owen values in tree-based ensembles by combining hierarchy-guided group aggregation with tree-aware dynamic programming. The exact algorithm computes Owen values without sampling under the path-dependent characteristic function, which approximates the conditional expectation, whereas the Monte Carlo algorithm provides a scalable approximation that is unbiased for any prespecified sampling budget and converges almost surely as the sampling budget increases. We also provide global importance measures and visualization tools for structured, multi-resolution explanations. The proposed algorithms and tools are collectively referred to as TreeOwen. Through simulation experiments, we demonstrate the numerical accuracy and substantial computational gains of TreeOwen. We illustrate its practical utility using immunotherapy metagenomic data, showing how microbial genera (groups) and species (features) contribute to patient recovery.

17
A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences

Motta, J. A.; Motta, M. d. M.; Fernandez, C.

2026-08-20 bioinformatics 10.64898/2026.08.16.745093 medRxiv
Top 0.6%
0.5%
Show abstract

In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.

18
pbcftools: parallel execution of bcftools for large variant call sets

Zhang, G.

2026-08-09 bioinformatics 10.64898/2026.08.03.742604 medRxiv
Top 0.7%
0.5%
Show abstract

Summarybcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the results by hand. We present pbcftools, a Perl wrapper that does this automatically: it splits the genome into chunks, runs an ordinary bcftools command on each in parallel, and reassembles the outputs by a method suited to the data type. Across Linux servers, Windows/WSL2 workstations and Apple laptops, with bcftools 1.21 to 1.24, parallel output was identical to serial output for every command tested. On 1000 Genomes Phase 3 data, operations writing compressed VCF ran 10.8 to 21.1 times faster with 32 cores and up to 35.4 times with 64, those writing text 3.7 to 12.8 times, and merging 100 VCF files 19.2 times. pbcftools also runs on LSF and Slurm clusters. Availability and implementationpbcftools is written in Perl (>= 5.16) and requires bcftools; local parallel execution also requires Perl module Parallel::ForkManager. It is released under the MIT license at https://github.com/zhangge-uc/pbcftools (DOI: 10.5281/zenodo.21780361).

19
SVPopEx: Population-Wide Visualization and Exploration of Structural Variants

Baker, M.; Bett, K.; Vargas, A.; Jin, L.

2026-08-14 bioinformatics 10.64898/2026.08.08.743609 medRxiv
Top 0.7%
0.4%
Show abstract

Structural variants (SVs) are large-scale genomic variants, which can disrupt important functional and regulatory elements, leading to genomic disorders in humans and playing important roles in domestication, disease resistance, and traits in plants. SVs are generated across populations of individuals and used for association studies, consisting of large datasets with thousands of genomic loci. Visualization of these SVs aids in understanding their genomic distribution, identifying patterns across affected or phenotypic groups, and assessing their proximity to other genomic regions of interest. A variety of tools exist for visualizing SVs, including linear genome browsers and graph-based methods; however, many do not offer intuitive or scalable representations of SVs across large populations. To address this, we present SVPopEx, an interactive tool for population-wide visualization and exploration of SVs. SVPopEx provides a unique and intuitive representation for insertions, deletions, inversions, duplications, and translocations in a linear genome-style browser. Novel features were developed to support comparisons across genomes within user-defined regions, including rendering SVs based on one or more samples and visualizing haplotypes. Use of the tool is demonstrated with SV datasets from Schistosoma mansoni and Lens culinaris. A task-based evaluation was conducted using SVPopEx and two other linear genome browsers, which demonstrated that SVPopEx excelled in (1) providing a clear representation of the SVs present and (2) supporting comparisons across genomes.

20
Correction of the cytosine deamination artifacts in FFPE-based sequencing experiments

Płonka, W.; Kostka, D.; Lalik, A.; Kurpas, M.; Dinh, K. N.; Sitkiewicz, M.; Kimmel, M.; Rzyman, W.; Jaksik, R.

2026-08-19 bioinformatics 10.64898/2026.08.11.744151 medRxiv
Top 0.8%
0.4%
Show abstract

Formalin-fixed, paraffin-embedded (FFPE) tissues remain an essential resource for molecular studies, yet formalin-induced cytosine deamination introduces characteristic C>T/G>A artifacts that compromise the accuracy of next-generation sequencing (NGS) analyses. Numerous computational methods and enzymatic DNA repair strategies have been proposed to reduce these artifacts, but no systematic comparison across tools and experimental conditions exists. Here, we evaluate the performance of seven computational approaches (SOBDetector, Ideafix, MicroSEC, FFPolish, DeepOmics FFPE/FFPE-PLUS, FFPErase) together with the NEBNext(R) FFPE DNA Repair Mix v2, a multi-enzyme repair system applied during DNA preparation. Using three independent datasets, one based on whole genome sequencing (CGCI-BL) and two on whole exome sequencing (TCGA-PC and SUT-LUAD, the latter containing enzymatically repaired samples), and matched fresh-frozen samples as the gold standard, we assess precision, sensitivity, and artifact reduction efficiency across all methods. We further examine the potential synergy between enzymatic repair and post-sequencing computational filtering. Our results provide practical guidelines for FFPE artifact correction and demonstrate that enzymatic treatment provides the best results, while among the computational methods, FFPErase offers the most robust reduction of cytosine deamination artifacts while maximizing the retention of true somatic variants. KEY MESSAGESO_LIFormalin fixation in FFPE samples introduces artifacts that can significantly affect the accuracy of NGS analyses. C_LIO_LIAmong the evaluated approaches, enzymatic repair using NEBNext(R) FFPE DNA Repair Mix v2 achieves the most effective reduction of sequencing artifacts. C_LIO_LIComputational methods vary in performance, with FFPErase showing the most robust balance between artifact removal and retention of true somatic variants. C_LIO_LICombining enzymatic repair with computational filtering did not lead to consistent improvements in performance across datasets. C_LI